SPIKE::GPU A SPIKE-based preconditioned GPU Solver for Sparse Linear Systems

نویسندگان

  • Ang Li
  • Andrew Seidl
  • Radu Serban
  • Dan Negrut
چکیده

This contribution outlines an approach that draws on general purpose graphics processing unit (GPGPU) computing to solve large linear systems. To methodology proposed relies on a SPIKE-based preconditioner with a Krylov-subspace method and has the following three stages: (i) row/column reordering for boosting diagonal dominance and reducing bandwidth; (ii) applying single precision truncated SPIKE on a dense banded matrix obtained after dropping small elements that fall outside a carefully selected bandwidth; and (iii) preconditioning within the BiCGStab(2) framework. The reordering strategy adopted aims at generating a narrow bandwidth dense matrix that is diagonally heavy. The results of several numerical experiments indicate that when used in conjunction with large dense banded matrices, the proposed approach is two to three times faster than the latest version of the MKL dense solver as soon as d > 0.23. When handling sparse systems, synthetic results obtained for large random matrices suggest that the proposed approach is three to five times faster than the PARDISO solver as long as the reordered matrices are close to being diagonally dominant. For a random set of smaller dimension application matrices, some out of the University of Florida Sparse Matrix Collection, the proposed approach is shown to be better or comparable in performance to PARDISO.

برای دانلود متن کامل این مقاله و بیش از 32 میلیون مقاله دیگر ابتدا ثبت نام کنید

ثبت نام

اگر عضو سایت هستید لطفا وارد حساب کاربری خود شوید

منابع مشابه

A Comparison of the Performance of SaP::GPU and Intel’s Math Kernel Library (MKL) for Solving Dense Banded Linear Systems

SaP::GPU is a solver developed in the Simulation Based Engineering Lab (SBEL) [1] to solve large banded and sparse linear systems on the GPU. This report contributes the performance comparison of the banded solver of SaP::GPU and Intel’s Math Kernel Library [2] on a large set of synthetic problems. The results of several numerical experiments indicate that when it is used in conjunction with la...

متن کامل

Architecting the finite element method pipeline for the GPU

The finite element method (FEM) is a widely employed numerical technique for approximating the solution of partial differential equations (PDEs) in various science and engineering applications. Many of these applications benefit from fast execution of the FEM pipeline. One way to accelerate the FEM pipeline is by exploiting advances in modern computational hardware, such as the many-core stream...

متن کامل

Concurrent number cruncher: a GPU implementation of a general sparse linear solver

A wide class of numerical methods needs to solve a linear system, where the matrix pattern of non-zero coefficients can be arbitrary. These problems can greatly benefit from highly multithreaded computational power and large memory bandwidth available on GPUs, especially since dedicated general purpose APIs such as CTM (AMD-ATI) and CUDA (NVIDIA) have appeared. CUDA even provides a BLAS impleme...

متن کامل

A Parallel Algebraic Multigrid Solver on Graphics Processing Units

The paper presents a multi-GPU implementation of the preconditioned conjugate gradient algorithm with an algebraic multigrid preconditioner (PCG-AMG) for an elliptic model problem on a 3D unstructured grid. An efficient parallel sparse matrix-vector multiplication scheme underlying the PCG-AMG algorithm is presented for the manycore GPU architecture. A performance comparison of the parallel sol...

متن کامل

Development of Krylov and AMG Linear Solvers for Large-Scale Sparse Matrices on GPUs

This research introduce our work on developing Krylov subspace and AMG solvers on NVIDIA GPUs. As SpMV is a crucial part for these iterative methods, SpMV algorithms for single GPU and multiple GPUs are implemented. A HEC matrix format and a communication mechanism are established. And also, a set of specific algorithms for solving preconditioned systems in parallel environments are designed, i...

متن کامل

ذخیره در منابع من


  با ذخیره ی این منبع در منابع من، دسترسی به آن را برای استفاده های بعدی آسان تر کنید

برای دانلود متن کامل این مقاله و بیش از 32 میلیون مقاله دیگر ابتدا ثبت نام کنید

ثبت نام

اگر عضو سایت هستید لطفا وارد حساب کاربری خود شوید

عنوان ژورنال:

دوره   شماره 

صفحات  -

تاریخ انتشار 2014